Skip to content

Reduce and harden MCP E2E coverage - #176

Merged
kentcdodds merged 2 commits into
mainfrom
cursor/reduce-mcp-e2e-flake-3f6c
Apr 15, 2026
Merged

kentcdodds merged 2 commits into
mainfrom
cursor/reduce-mcp-e2e-flake-3f6c

Conversation

@kentcdodds

@kentcdodds kentcdodds commented Apr 15, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • shrink the MCP E2E suite to a tiny transport smoke file with an explicit warning against casually adding cases there
  • move detailed saved-app capability assertions into fast node-level tests beside the implementation
  • increase MCP E2E startup/test timeouts so cold Wrangler + OAuth runs get enough headroom

Testing

  • npx vitest run --project node-unit packages/worker/src/mcp/capabilities/apps/ui-save-app.node.test.ts packages/worker/src/mcp/capabilities/apps/saved-app-capabilities.node.test.ts
  • npm run test:mcp
  • npm run typecheck
Open in Web Open in Cursor 

Summary by CodeRabbit

Release Notes

  • Documentation

    • Updated testing guidance for MCP capability development, recommending colocated unit tests with selective E2E transport validation.
  • Tests

    • Added comprehensive test coverage for saved app capabilities including creation, updates, storage export, and execution.
    • Streamlined MCP E2E test suite to focus on critical smoke journeys.
    • Improved test infrastructure timeouts for environment-specific configurations.

Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
@coderabbitai

coderabbitai Bot commented Apr 15, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@cursor[bot] has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 43 minutes and 10 seconds before requesting another review.

Your organization is not enrolled in usage-based pricing. Contact your admin to enable usage-based pricing to continue reviews beyond the rate limit, or try again in 43 minutes and 10 seconds.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: a0671367-7e73-4295-aa3d-569d7f43acb1

📥 Commits

Reviewing files that changed from the base of the PR and between f8289da and e404d49.

📒 Files selected for processing (1)
  • packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts
📝 Walkthrough

Walkthrough

This pull request reorganizes MCP capability testing strategy by shifting from comprehensive E2E coverage per capability to focused unit tests colocated with implementation. Documentation is updated to reflect this approach, new focused .node.test.ts files are added for saved-app capabilities, the E2E suite is simplified to smoke tests, and test timeouts are adjusted.

Changes

Cohort / File(s) Summary
Testing Documentation
docs/contributing/adding-capabilities.md, docs/contributing/end-to-end-testing.md, docs/contributing/testing-principles.md
Updated guidance to recommend focused *.node.test.ts and *.workers.test.ts tests adjacent to implementation for MCP capabilities, reserving *.mcp-e2e.test.ts for smoke tests and behaviors dependent on real MCP transport, OAuth flow, or session wiring.
Saved App Capability Tests
packages/worker/src/mcp/capabilities/apps/saved-app-capabilities.node.test.ts, packages/worker/src/mcp/capabilities/apps/ui-save-app.node.test.ts
New focused Vitest suites for saved-app capability handlers with mocked dependencies, validating app_server_exec, app_storage_export, ui_load_app_source, and uiSaveAppCapability observable behavior and update scenarios.
E2E Test Suite Restructuring
packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts
Substantially simplified from comprehensive per-capability coverage to minimal smoke journeys; removed deep capability-specific assertions and test cases while retaining focused validation of open_generated_ui output and saved-app reopening behavior.
Test Configuration
tools/mcp-test-support.ts, vitest.mcp-e2e.config.ts
Updated server readiness and MCP E2E timeout values to be environment-dependent, increasing non-CI local timeouts from 10–25 seconds to 45–60 seconds for better stability.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Poem

🐰 Hop, hop—less E2E broad strokes, more focused node tests take flight!
Smoke journeys trim the suite, while focused tests keep behaviors tight.
Doc updates guide the path, capabilities tested right, right, right! 🚀

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main objective of the changeset: reducing MCP E2E coverage and hardening it through focused smoke tests.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch cursor/reduce-mcp-e2e-flake-3f6c

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@kentcdodds
kentcdodds marked this pull request as ready for review April 15, 2026 04:07
@github-actions

github-actions Bot commented Apr 15, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Preview deployed: https://kody-pr-176.kentcdodds.workers.dev

Worker: kody-pr-176
D1: kody-pr-176-db
KV: kody-pr-176-oauth-kv

Mocks:

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts`:
- Line 266: The test currently sets a hardcoded per-test timeout of 45_000 (the
trailing "45_000" literal at the end of the test call) which overrides the CI
config (vitest.mcp-e2e.config.ts); remove the per-test timeout so the test
inherits the config timeout, or replace the literal with an environment-aware
value (e.g., read a MAX_TEST_TIMEOUT env var or use process.env.CI to pick
120_000) and apply that instead in the same test invocation; update the test
invocation in packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts where the
45_000 literal appears.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e5fee5dc-da2d-4be1-a588-e54e2036e991

📥 Commits

Reviewing files that changed from the base of the PR and between 2d0f88e and f8289da.

📒 Files selected for processing (8)
  • docs/contributing/adding-capabilities.md
  • docs/contributing/end-to-end-testing.md
  • docs/contributing/testing-principles.md
  • packages/worker/src/mcp/capabilities/apps/saved-app-capabilities.node.test.ts
  • packages/worker/src/mcp/capabilities/apps/ui-save-app.node.test.ts
  • packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts
  • tools/mcp-test-support.ts
  • vitest.mcp-e2e.config.ts

Comment thread packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

Bugbot Autofix prepared a fix for the issue found in the latest run.

  • ✅ Fixed: Per-test timeout overrides CI-aware global timeout
    • Removed the per-test timeout override so the CI-aware global timeout applies to the consolidated MCP smoke test.
Preview (e404d4905a)
diff --git a/docs/contributing/adding-capabilities.md b/docs/contributing/adding-capabilities.md
--- a/docs/contributing/adding-capabilities.md
+++ b/docs/contributing/adding-capabilities.md
@@ -219,8 +219,11 @@
    `capabilities: [..., yourCapability]`.
 5. If the domain uses `index.ts`, ensure it still exports `domain` /
    `codingCapabilities`-style aliases as needed for local imports.
-6. Add or update tests in `packages/worker/src/mcp/*.mcp-e2e.test.ts` for
-   MCP-visible behavior.
+6. Add or update focused `*.node.test.ts` or `*.workers.test.ts` coverage beside
+   the implementation for most MCP-visible behavior. Touch
+   `packages/worker/src/mcp/*.mcp-e2e.test.ts` only when the behavior truly
+   depends on the real MCP transport, OAuth handshake, or saved-app session
+   wiring.
 
 Example (assuming `example` exists in `capabilityDomainNames`):
 
@@ -317,8 +320,9 @@
   ranked results
 - use `execute` to confirm the capability runs correctly
 
-Prefer E2E tests in `packages/worker/src/mcp/*.mcp-e2e.test.ts` for the real MCP
-contract.
+Prefer `*.node.test.ts` and `*.workers.test.ts` for capability behavior. Reserve
+`packages/worker/src/mcp/*.mcp-e2e.test.ts` for a very small number of real MCP
+contract smoke tests.
 
 Registry invariants (duplicate capability names, domain/capability mismatches,
 duplicate domain registration) are covered in

diff --git a/docs/contributing/end-to-end-testing.md b/docs/contributing/end-to-end-testing.md
--- a/docs/contributing/end-to-end-testing.md
+++ b/docs/contributing/end-to-end-testing.md
@@ -31,6 +31,10 @@
   cases.
 - If a bug is unlikely to show up again, do not add an E2E test just to lock in
   the fix.
+- For MCP specifically, treat `*.mcp-e2e.test.ts` as a tiny transport smoke
+  suite. Do not add capability-by-capability coverage there unless the failure
+  mode depends on the real MCP HTTP transport, OAuth flow, or saved-app session
+  wiring.
 
 ## Structure and style
 
@@ -108,6 +112,10 @@
 - `npx playwright test`
 - `npx playwright test e2e/login.spec.ts`
 
+For MCP capability work, prefer `*.node.test.ts` or `*.workers.test.ts` beside
+the implementation and keep `npm run test:mcp` limited to a couple of
+high-signal smoke journeys.
+
 If `packages/worker/.env` is missing, the E2E server startup path copies
 `packages/worker/.env.example` to `packages/worker/.env` before Wrangler starts.
 

diff --git a/docs/contributing/testing-principles.md b/docs/contributing/testing-principles.md
--- a/docs/contributing/testing-principles.md
+++ b/docs/contributing/testing-principles.md
@@ -30,6 +30,9 @@
   tests.
 - Prefer fast unit tests for server logic; keep e2e tests focused on a very
   small number of important happy-path journeys.
+- Treat `packages/worker/src/mcp/*.mcp-e2e.test.ts` as a tiny MCP transport
+  smoke suite. Do not add capability-specific cases there unless they require
+  the real MCP HTTP transport, OAuth flow, and saved-app session wiring.
 - Prefer asserting intermediate states inside the broader workflow that causes
   them rather than adding isolated tests that only check an incidental loading
   or transition state.

diff --git a/packages/worker/src/mcp/capabilities/apps/saved-app-capabilities.node.test.ts b/packages/worker/src/mcp/capabilities/apps/saved-app-capabilities.node.test.ts
new file mode 100644
--- /dev/null
+++ b/packages/worker/src/mcp/capabilities/apps/saved-app-capabilities.node.test.ts
@@ -1,0 +1,205 @@
+import { expect, test, vi } from 'vitest'
+import { createMcpCallerContext } from '#mcp/context.ts'
+
+const mockModule = vi.hoisted(() => ({
+	syncSavedAppRunnerFromDb: vi.fn(),
+	execSavedAppRunnerServer: vi.fn(),
+	exportSavedAppRunnerStorage: vi.fn(),
+	getUiArtifactById: vi.fn(),
+}))
+
+vi.mock('#mcp/app-runner.ts', () => ({
+	syncSavedAppRunnerFromDb: (...args: Array<unknown>) =>
+		mockModule.syncSavedAppRunnerFromDb(...args),
+	execSavedAppRunnerServer: (...args: Array<unknown>) =>
+		mockModule.execSavedAppRunnerServer(...args),
+	exportSavedAppRunnerStorage: (...args: Array<unknown>) =>
+		mockModule.exportSavedAppRunnerStorage(...args),
+}))
+
+vi.mock('#mcp/ui-artifacts-repo.ts', () => ({
+	getUiArtifactById: (...args: Array<unknown>) =>
+		mockModule.getUiArtifactById(...args),
+}))
+
+const { appServerExecCapability } = await import('./app-server-exec.ts')
+const { appStorageExportCapability } = await import('./app-storage-export.ts')
+const { uiLoadAppSourceCapability } = await import('./ui-load-app-source.ts')
+
+test('app_server_exec syncs the runner and normalizes the runner response', async () => {
+	mockModule.syncSavedAppRunnerFromDb.mockReset()
+	mockModule.execSavedAppRunnerServer.mockReset()
+	mockModule.exportSavedAppRunnerStorage.mockReset()
+	mockModule.getUiArtifactById.mockReset()
+
+	mockModule.syncSavedAppRunnerFromDb.mockResolvedValueOnce({
+		id: 'app-1',
+	})
+	mockModule.execSavedAppRunnerServer.mockResolvedValueOnce({
+		appId: 'app-1',
+		facetName: 'main',
+		result: { count: 3 },
+	})
+
+	const result = await appServerExecCapability.handler(
+		{
+			app_id: 'app-1',
+			code: 'return await app.call("incrementBy", params.amount ?? 1)',
+			params: { amount: 3 },
+		},
+		{
+			env: {} as Env,
+			callerContext: createMcpCallerContext({
+				baseUrl: 'https://heykody.dev',
+				user: { userId: 'user-1', email: 'user@example.com' },
+			}),
+		},
+	)
+
+	expect(mockModule.syncSavedAppRunnerFromDb).toHaveBeenCalledWith({
+		env: {},
+		appId: 'app-1',
+		userId: 'user-1',
+		baseUrl: 'https://heykody.dev',
+	})
+	expect(mockModule.execSavedAppRunnerServer).toHaveBeenCalledWith({
+		env: {},
+		appId: 'app-1',
+		facetName: 'main',
+		code: 'return await app.call("incrementBy", params.amount ?? 1)',
+		params: { amount: 3 },
+	})
+	expect(result).toEqual({
+		ok: true,
+		app_id: 'app-1',
+		facet_name: 'main',
+		result: { count: 3 },
+	})
+})
+
+test('app_storage_export forwards pagination options to the runner export helper', async () => {
+	mockModule.syncSavedAppRunnerFromDb.mockReset()
+	mockModule.execSavedAppRunnerServer.mockReset()
+	mockModule.exportSavedAppRunnerStorage.mockReset()
+	mockModule.getUiArtifactById.mockReset()
+
+	mockModule.syncSavedAppRunnerFromDb.mockResolvedValueOnce({
+		id: 'app-1',
+	})
+	mockModule.exportSavedAppRunnerStorage.mockResolvedValueOnce({
+		appId: 'app-1',
+		facetName: 'analytics',
+		export: {
+			entries: [{ key: 'count', value: 3 }],
+			estimatedBytes: 128,
+			truncated: true,
+			nextStartAfter: 'count',
+			pageSize: 1,
+		},
+	})
+
+	const result = await appStorageExportCapability.handler(
+		{
+			app_id: 'app-1',
+			facet_name: 'analytics',
+			page_size: 1,
+			start_after: 'count',
+		},
+		{
+			env: {} as Env,
+			callerContext: createMcpCallerContext({
+				baseUrl: 'https://heykody.dev',
+				user: { userId: 'user-1', email: 'user@example.com' },
+			}),
+		},
+	)
+
+	expect(mockModule.syncSavedAppRunnerFromDb).toHaveBeenCalledWith({
+		env: {},
+		appId: 'app-1',
+		userId: 'user-1',
+		baseUrl: 'https://heykody.dev',
+	})
+	expect(mockModule.exportSavedAppRunnerStorage).toHaveBeenCalledWith({
+		env: {},
+		appId: 'app-1',
+		facetName: 'analytics',
+		pageSize: 1,
+		startAfter: 'count',
+	})
+	expect(result).toEqual({
+		ok: true,
+		app_id: 'app-1',
+		facet_name: 'analytics',
+		export: {
+			entries: [{ key: 'count', value: 3 }],
+			estimatedBytes: 128,
+			truncated: true,
+			nextStartAfter: 'count',
+			pageSize: 1,
+		},
+	})
+})
+
+test('ui_load_app_source returns saved source for the authenticated user', async () => {
+	mockModule.syncSavedAppRunnerFromDb.mockReset()
+	mockModule.execSavedAppRunnerServer.mockReset()
+	mockModule.exportSavedAppRunnerStorage.mockReset()
+	mockModule.getUiArtifactById.mockReset()
+
+	mockModule.getUiArtifactById.mockResolvedValueOnce({
+		id: 'app-1',
+		title: 'Patchable App',
+		description: 'Saved app source',
+		clientCode: '<main><h1>Saved</h1></main>',
+		serverCode:
+			'import { DurableObject } from "cloudflare:workers"; export class App extends DurableObject {}',
+		serverCodeId: 'server-code-v1',
+		parameters: JSON.stringify([
+			{
+				name: 'team',
+				description: 'Team slug',
+				type: 'string',
+				required: true,
+			},
+		]),
+		hidden: true,
+	})
+
+	const result = await uiLoadAppSourceCapability.handler(
+		{
+			app_id: 'app-1',
+		},
+		{
+			env: { APP_DB: {} } as Env,
+			callerContext: createMcpCallerContext({
+				baseUrl: 'https://heykody.dev',
+				user: { userId: 'user-1', email: 'user@example.com' },
+			}),
+		},
+	)
+
+	expect(mockModule.getUiArtifactById).toHaveBeenCalledWith(
+		{},
+		'user-1',
+		'app-1',
+	)
+	expect(result).toEqual({
+		app_id: 'app-1',
+		title: 'Patchable App',
+		description: 'Saved app source',
+		client_code: '<main><h1>Saved</h1></main>',
+		server_code:
+			'import { DurableObject } from "cloudflare:workers"; export class App extends DurableObject {}',
+		server_code_id: 'server-code-v1',
+		parameters: [
+			{
+				name: 'team',
+				description: 'Team slug',
+				type: 'string',
+				required: true,
+			},
+		],
+		hidden: true,
+	})
+})

diff --git a/packages/worker/src/mcp/capabilities/apps/ui-save-app.node.test.ts b/packages/worker/src/mcp/capabilities/apps/ui-save-app.node.test.ts
new file mode 100644
--- /dev/null
+++ b/packages/worker/src/mcp/capabilities/apps/ui-save-app.node.test.ts
@@ -1,0 +1,249 @@
+import { expect, test, vi } from 'vitest'
+import { createMcpCallerContext } from '#mcp/context.ts'
+
+const mockModule = vi.hoisted(() => ({
+	getUiArtifactById: vi.fn(),
+	updateUiArtifact: vi.fn(),
+	insertUiArtifact: vi.fn(),
+	deleteUiArtifact: vi.fn(),
+	configureSavedAppRunner: vi.fn(),
+	deleteSavedAppRunner: vi.fn(),
+	upsertUiArtifactVector: vi.fn(),
+	deleteUiArtifactVector: vi.fn(),
+}))
+
+vi.mock('#mcp/ui-artifacts-repo.ts', () => ({
+	getUiArtifactById: (...args: Array<unknown>) =>
+		mockModule.getUiArtifactById(...args),
+	updateUiArtifact: (...args: Array<unknown>) =>
+		mockModule.updateUiArtifact(...args),
+	insertUiArtifact: (...args: Array<unknown>) =>
+		mockModule.insertUiArtifact(...args),
+	deleteUiArtifact: (...args: Array<unknown>) =>
+		mockModule.deleteUiArtifact(...args),
+}))
+
+vi.mock('#mcp/app-runner.ts', () => ({
+	configureSavedAppRunner: (...args: Array<unknown>) =>
+		mockModule.configureSavedAppRunner(...args),
+	deleteSavedAppRunner: (...args: Array<unknown>) =>
+		mockModule.deleteSavedAppRunner(...args),
+}))
+
+vi.mock('#mcp/ui-artifacts-vectorize.ts', () => ({
+	upsertUiArtifactVector: (...args: Array<unknown>) =>
+		mockModule.upsertUiArtifactVector(...args),
+	deleteUiArtifactVector: (...args: Array<unknown>) =>
+		mockModule.deleteUiArtifactVector(...args),
+}))
+
+const { uiSaveAppCapability } = await import('./ui-save-app.ts')
+
+test('ui_save_app updates preserve backend code unless the caller clears or replaces it', async () => {
+	mockModule.getUiArtifactById.mockReset()
+	mockModule.updateUiArtifact.mockReset()
+	mockModule.insertUiArtifact.mockReset()
+	mockModule.deleteUiArtifact.mockReset()
+	mockModule.configureSavedAppRunner.mockReset()
+	mockModule.deleteSavedAppRunner.mockReset()
+	mockModule.upsertUiArtifactVector.mockReset()
+	mockModule.deleteUiArtifactVector.mockReset()
+
+	const initialServerCode =
+		'import { DurableObject } from "cloudflare:workers"; export class App extends DurableObject { async readVersion() { return "v1" } }'
+	const replacementServerCode =
+		'import { DurableObject } from "cloudflare:workers"; export class App extends DurableObject { async readVersion() { return "v2" } }'
+
+	let currentApp = {
+		id: 'app-1',
+		user_id: 'user-1',
+		title: 'Patchable App',
+		description: 'Saved app used to verify partial ui_save_app updates.',
+		clientCode: '<main><h1>Patchable v1</h1></main>',
+		serverCode: initialServerCode,
+		serverCodeId: 'server-code-v1',
+		parameters: JSON.stringify([
+			{
+				name: 'team',
+				description: 'Team slug',
+				type: 'string',
+				required: true,
+			},
+		]),
+		hidden: true,
+	}
+
+	mockModule.getUiArtifactById.mockImplementation(async () => ({
+		...currentApp,
+	}))
+	mockModule.updateUiArtifact.mockImplementation(
+		async (
+			_db: unknown,
+			_userId: string,
+			_appId: string,
+			updates: Record<string, unknown>,
+		) => {
+			currentApp = {
+				...currentApp,
+				...(updates['title'] !== undefined
+					? { title: updates['title'] as string }
+					: {}),
+				...(updates['description'] !== undefined
+					? { description: updates['description'] as string }
+					: {}),
+				...(updates['clientCode'] !== undefined
+					? { clientCode: updates['clientCode'] as string }
+					: {}),
+				...(updates['hidden'] !== undefined
+					? { hidden: updates['hidden'] as boolean }
+					: {}),
+				...(updates['parameters'] !== undefined
+					? { parameters: updates['parameters'] as string | null }
+					: {}),
+				...(updates['serverCode'] !== undefined
+					? { serverCode: updates['serverCode'] as string | null }
+					: {}),
+				...(updates['serverCodeId'] !== undefined
+					? { serverCodeId: updates['serverCodeId'] as string }
+					: {}),
+			}
+			return { ...currentApp }
+		},
+	)
+
+	const randomUuidSpy = vi.spyOn(crypto, 'randomUUID')
+	randomUuidSpy.mockReturnValueOnce('server-code-v2')
+	randomUuidSpy.mockReturnValueOnce('server-code-v3')
+
+	try {
+		const callerContext = createMcpCallerContext({
+			baseUrl: 'https://heykody.dev',
+			user: { userId: 'user-1', email: 'user@example.com' },
+		})
+
+		const preservedResult = await uiSaveAppCapability.handler(
+			{
+				app_id: 'app-1',
+				clientCode: '<main><h1>Patchable v2</h1></main>',
+			},
+			{
+				env: { APP_DB: {} } as Env,
+				callerContext,
+			},
+		)
+		expect(preservedResult).toEqual({
+			app_id: 'app-1',
+			server_code_id: 'server-code-v1',
+			has_server_code: true,
+			hosted_url: 'https://heykody.dev/ui/app-1',
+			parameters: [
+				{
+					name: 'team',
+					description: 'Team slug',
+					type: 'string',
+					required: true,
+				},
+			],
+			hidden: true,
+		})
+		expect(mockModule.updateUiArtifact.mock.calls[0]?.[3]).toEqual({
+			title: undefined,
+			description: undefined,
+			clientCode: '<main><h1>Patchable v2</h1></main>',
+			hidden: undefined,
+		})
+		expect(mockModule.configureSavedAppRunner.mock.calls[0]?.[0]).toEqual(
+			expect.objectContaining({
+				appId: 'app-1',
+				serverCode: initialServerCode,
+				serverCodeId: 'server-code-v1',
+			}),
+		)
+
+		const clearedResult = await uiSaveAppCapability.handler(
+			{
+				app_id: 'app-1',
+				serverCode: null,
+			},
+			{
+				env: { APP_DB: {} } as Env,
+				callerContext,
+			},
+		)
+		expect(clearedResult).toEqual({
+			app_id: 'app-1',
+			server_code_id: 'server-code-v2',
+			has_server_code: false,
+			hosted_url: 'https://heykody.dev/ui/app-1',
+			parameters: [
+				{
+					name: 'team',
+					description: 'Team slug',
+					type: 'string',
+					required: true,
+				},
+			],
+			hidden: true,
+		})
+		expect(mockModule.updateUiArtifact.mock.calls[1]?.[3]).toEqual({
+			title: undefined,
+			description: undefined,
+			clientCode: undefined,
+			hidden: undefined,
+			serverCode: null,
+			serverCodeId: 'server-code-v2',
+		})
+		expect(mockModule.configureSavedAppRunner.mock.calls[1]?.[0]).toEqual(
+			expect.objectContaining({
+				appId: 'app-1',
+				serverCode: null,
+				serverCodeId: 'server-code-v2',
+			}),
+		)
+
+		const replacedResult = await uiSaveAppCapability.handler(
+			{
+				app_id: 'app-1',
+				serverCode: replacementServerCode,
+			},
+			{
+				env: { APP_DB: {} } as Env,
+				callerContext,
+			},
+		)
+		expect(replacedResult).toEqual({
+			app_id: 'app-1',
+			server_code_id: 'server-code-v3',
+			has_server_code: true,
+			hosted_url: 'https://heykody.dev/ui/app-1',
+			parameters: [
+				{
+					name: 'team',
+					description: 'Team slug',
+					type: 'string',
+					required: true,
+				},
+			],
+			hidden: true,
+		})
+		expect(mockModule.updateUiArtifact.mock.calls[2]?.[3]).toEqual({
+			title: undefined,
+			description: undefined,
+			clientCode: undefined,
+			hidden: undefined,
+			serverCode: replacementServerCode,
+			serverCodeId: 'server-code-v3',
+		})
+		expect(mockModule.configureSavedAppRunner.mock.calls[2]?.[0]).toEqual(
+			expect.objectContaining({
+				appId: 'app-1',
+				serverCode: replacementServerCode,
+				serverCodeId: 'server-code-v3',
+			}),
+		)
+		expect(randomUuidSpy).toHaveBeenCalledTimes(2)
+		expect(mockModule.deleteUiArtifactVector).toHaveBeenCalledTimes(3)
+	} finally {
+		randomUuidSpy.mockRestore()
+	}
+})

diff --git a/packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts b/packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts
--- a/packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts
+++ b/packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts
@@ -1,12 +1,20 @@
 import { expect, test } from 'vitest'
 import { type CallToolResult } from '@modelcontextprotocol/sdk/types.js'
-import { setTimeout as delay } from 'node:timers/promises'
 import {
 	createMcpClient,
 	createTestDatabase,
 	startDevServer,
 } from '../../../../tools/mcp-test-support.ts'
 
+/**
+ * MCP E2E is intentionally tiny.
+ *
+ * Do not add cases here unless the thing being tested genuinely requires the
+ * real MCP HTTP transport, OAuth flow, and saved-app session wiring all at
+ * once. Most capability behavior belongs in faster node/workers tests beside
+ * the implementation. Keep this file to a couple of smoke journeys.
+ */
+
 test('mcp endpoint requires OAuth bearer auth', async () => {
 	await using database = await createTestDatabase()
 	await using server = await startDevServer(database.persistDir)
@@ -25,7 +33,7 @@
 	)
 })
 
-test('authenticated mcp client can list tools, execute codemode, and search memories', async () => {
+test('authenticated MCP smoke covers core tools, inline UI, and saved app backends', async () => {
 	await using database = await createTestDatabase()
 	await using server = await startDevServer(database.persistDir)
 	await using mcpClient = await createMcpClient(server.origin, database.user)
@@ -55,85 +63,30 @@
 		name: 'execute',
 		arguments: {
 			code: `async () => {
-				return await codemode.meta_memory_upsert({
-					subject: 'User prefers npm over pnpm',
-					summary: 'Always use npm commands in this repository.',
-					category: 'preference',
-					tags: ['package-manager', 'repo-workflow'],
-					source_uris: [
-						'https://docs.npmjs.com/cli/v11/commands/npm-install',
-						'https://github.com/kentcdodds/kody/blob/main/AGENTS.md',
-					],
-					verified_by_agent: true,
-					verification_reference: 'verify-search-fallback-1',
-				})
-			}`,
+					return await codemode.meta_memory_upsert({
+						subject: 'User prefers npm over pnpm',
+						summary: 'Always use npm commands in this repository.',
+						category: 'preference',
+						tags: ['package-manager', 'repo-workflow'],
+						source_uris: [
+							'https://docs.npmjs.com/cli/v11/commands/npm-install',
+							'https://github.com/kentcdodds/kody/blob/main/AGENTS.md',
+						],
+						verified_by_agent: true,
+						verification_reference: 'verify-search-fallback-1',
+					})
+				}`,
 		},
 	})
 	const upsertStructured = (upsertResult as CallToolResult).structuredContent as
-		| { result?: { memory?: { id?: string; source_uris?: Array<string> } } }
+		| { result?: { memory?: { id?: string } } }
 		| undefined
 	expect(typeof upsertStructured?.result?.memory?.id).toBe('string')
-	expect(upsertStructured?.result?.memory?.source_uris).toEqual([
-		'https://docs.npmjs.com/cli/v11/commands/npm-install',
-		'https://github.com/kentcdodds/kody/blob/main/AGENTS.md',
-	])
 
-	const memoryId = upsertStructured?.result?.memory?.id
-	if (!memoryId) {
-		throw new Error('missing memoryId')
-	}
-
-	const getResult = await mcpClient.client.callTool({
-		name: 'execute',
-		arguments: {
-			code: `async () => {
-				return await codemode.meta_memory_get({
-					memory_id: ${JSON.stringify(memoryId)},
-				})
-			}`,
-		},
-	})
-	const getStructured = (getResult as CallToolResult).structuredContent as
-		| { result?: { source_uris?: Array<string> } }
-		| undefined
-	expect(getStructured?.result?.source_uris).toEqual([
-		'https://docs.npmjs.com/cli/v11/commands/npm-install',
-		'https://github.com/kentcdodds/kody/blob/main/AGENTS.md',
-	])
-
-	const memoryCapabilitySearchResult = await mcpClient.client.callTool({
-		name: 'execute',
-		arguments: {
-			code: `async () => {
-				return await codemode.meta_memory_search({
-					query: 'npm over pnpm',
-					limit: 3,
-				})
-			}`,
-		},
-	})
-	const memoryCapabilitySearchStructured = (
-		memoryCapabilitySearchResult as CallToolResult
-	).structuredContent as
-		| {
-				result?: {
-					matches?: Array<{ source_uris?: Array<string> }>
-				}
-		  }
-		| undefined
-	expect(
-		memoryCapabilitySearchStructured?.result?.matches?.[0]?.source_uris,
-	).toEqual([
-		'https://docs.npmjs.com/cli/v11/commands/npm-install',
-		'https://github.com/kentcdodds/kody/blob/main/AGENTS.md',
-	])
-
-	const query = 'npm over pnpm'
 	const memorySearchResult = await mcpClient.client.callTool({
 		name: 'search',
 		arguments: {
-			query,
+			query: 'npm over pnpm',
 		},
 	})
 	const memorySearchStructured = (memorySearchResult as CallToolResult)
@@ -147,8 +100,9 @@
 				}
 		  }
 		| undefined
-
-	expect(memorySearchStructured?.result?.memories?.retrievalQuery).toBe(query)
+	expect(memorySearchStructured?.result?.memories?.retrievalQuery).toBe(
+		'npm over pnpm',
+	)
 	expect(memorySearchStructured?.result?.memories?.surfaced).toEqual(
 		expect.arrayContaining([
 			expect.objectContaining({
@@ -156,20 +110,15 @@
 			}),
 		]),
 	)
-})
 
-test('authenticated mcp client can open generated ui and reopen a saved app', async () => {
-	await using database = await createTestDatabase()
-	await using server = await startDevServer(database.persistDir)
-	await using mcpClient = await createMcpClient(server.origin, database.user)
-
-	const inlineResult = await mcpClient.client.callTool({
+	const inlineUiResult = await mcpClient.client.callTool({
 		name: 'open_generated_ui',
 		arguments: {
 			code: '<main><h1>Storage Context</h1></main>',
 		},
 	})
-	const inlineStructured = (inlineResult as CallToolResult).structuredContent as
+	const inlineUiStructured = (inlineUiResult as CallToolResult)
+		.structuredContent as
 		| {
 				renderSource?: string
 				appSession?: {
@@ -178,104 +127,58 @@
 				} | null
 		  }
 		| undefined
-	const executeEndpoint = inlineStructured?.appSession?.endpoints?.execute
-	const executeToken = inlineStructured?.appSession?.token
-	expect(inlineStructured?.renderSource).toBe('inline_code')
-	expect(typeof executeEndpoint).toBe('string')
-	expect(typeof executeToken).toBe('string')
-
-	const setValueResponse = await fetch(executeEndpoint!, {
-		method: 'POST',
-		headers: {
-			Authorization: `Bearer ${executeToken}`,
-			'Content-Type': 'application/json',
-			Accept: 'application/json',
-		},
-		body: JSON.stringify({
-			code: `async () => {
-				await codemode.value_set({
-					name: 'example',
-					value: 'value',
-					scope: 'session',
-				})
-				return { ok: true }
-			}`,
-		}),
-	})
-	expect(setValueResponse.ok).toBe(true)
-	const setValuePayload = (await setValueResponse.json()) as {
-		ok?: boolean
-		result?: { ok?: boolean }
-	}
-	expect(setValuePayload.ok).toBe(true)
-	expect(setValuePayload.result).toEqual({ ok: true })
-
-	// Repeated POSTs to the generated UI execute endpoint can hit Wrangler local
-	// dev proxy restarts mid-request in CI. The storage-backed execute behavior is
-	// covered by focused unit/workers tests, so this E2E keeps the generated UI
-	// flow coverage without asserting a second POST round-trip here.
-	await delay(2500)
-
-	// The generated UI runtime executes out-of-band HTTP requests with its own app
-	// session. Reconnect the MCP client before resuming tool calls so the test
-	// exercises a fresh MCP session after that browser-style interaction.
-	await using resumedMcpClient = await createMcpClient(
-		server.origin,
-		database.user,
+	expect(inlineUiStructured?.renderSource).toBe('inline_code')
+	expect(typeof inlineUiStructured?.appSession?.token).toBe('string')
+	expect(typeof inlineUiStructured?.appSession?.endpoints?.execute).toBe(
+		'string',
 	)
 
-	const saveResult = await resumedMcpClient.client.callTool({
+	const saveResult = await mcpClient.client.callTool({
 		name: 'execute',
 		arguments: {
 			code: `async () => {
-				return await codemode.ui_save_app({
-					title: 'Persistent UI',
-					description: 'Saved from test',
-					clientCode: '<main><h1>Saved</h1></main>',
-					hidden: false,
-				})
-			}`,
+					return await codemode.ui_save_app({
+						title: 'Facet Counter',
+						description: 'Saved app with a backend facet counter',
+						clientCode: '<main><h1>Facet Counter</h1></main>',
+						serverCode: \`
+							import { DurableObject } from 'cloudflare:workers'
+
+							export class App extends DurableObject {
+								async fetch(request) {
+									const url = new URL(request.url)
+									if (url.pathname !== '/api/counter') {
+										return new Response('Not found', { status: 404 })
+									}
+									const current = (await this.ctx.storage.get('count')) ?? 0
+									const next = Number(current) + 1
+									await this.ctx.storage.put('count', next)
+									return Response.json({ count: next })
+								}
+							}
+						\`,
+						hidden: true,
+					})
+				}`,
 		},
 	})
 	const saveStructured = (saveResult as CallToolResult).structuredContent as
-		| { result?: { app_id?: string } }
... diff truncated: showing 800 of 1341 lines

You can send follow-ups to the cloud agent here.

Reviewed by Cursor Bugbot for commit f8289da. Configure here.

Comment thread packages/worker/src/mcp/mcp-server.mcp-e2e.test.ts Outdated
Co-authored-by: Kent C. Dodds <me+github@kentcdodds.com>
@kentcdodds
kentcdodds merged commit 2eec1f5 into main Apr 15, 2026
9 checks passed
@kentcdodds
kentcdodds deleted the cursor/reduce-mcp-e2e-flake-3f6c branch April 17, 2026 01:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants